Search CORE

12 research outputs found

Női ÁVH-s koncepciós perek : az "ávós" nők jelenléte a kommunista igazságszolgáltatásban

Author: Lakatos Dorina
Publication venue: Szegedi Tudományegyetem Eötvös Loránd Kollégium
Publication date: 01/01/2022
Field of study

University of Szeged

Női ÁVH-s koncepciós perek

Author: Lakatos Dorina
Publication venue: Szegedi Tudományegyetem Eötvös Loránd Kollégium
Publication date: 01/01/2022
Field of study

University of Szeged

Data Augmentation for Machine Translation via Dependency Subtree Swapping

Author: Barta Botond
Lakatos Dorina Petra
Nagy Attila
Nanys Patrick
Ács Judit
Publication venue
Publication date: 13/07/2023
Field of study

We present a generic framework for data augmentation via dependency subtree swapping that is applicable to machine translation. We extract corresponding subtrees from the dependency parse trees of the source and target sentences and swap these across bisentences to create augmented samples. We perform thorough filtering based on graphbased similarities of the dependency trees and additional heuristics to ensure that extracted subtrees correspond to the same meaning. We conduct resource-constrained experiments on 4 language pairs in both directions using the IWSLT text translation datasets and the Hunglish2 corpus. The results demonstrate consistent improvements in BLEU score over our baseline models in 3 out of 4 language pairs. Our code is available on GitHub

arXiv.org e-Print Archive

Data augmentation for machine translation via dependency subtree swapping

Author: Barta Botond
Lakatos Dorina Petra
Nagy Attila
Nanys Patrick
Ács Judit
Publication venue
Publication date: 01/01/2023
Field of study

University of Szeged

HunSum-1 : an abstractive summarization dataset for Hungarian

Author: Barta Botond
Lakatos Dorina
Nagy Attila
Nyist Milán Konor
Ács Judit
Publication venue
Publication date: 01/01/2023
Field of study

We introduce HunSum-1 : a dataset for Hungarian abstractive summarization, consisting of 1.14M news articles. The dataset is built by collecting, cleaning and deduplicating data from 9 major Hungarian news sites through CommonCrawl. Using this dataset, we build abstractive summarizer models based on huBERT and mT5. We demonstrate the value of the created dataset by performing a quantitative and qualitative analysis on the models’ results. The HunSum-1 dataset, all models used in our experiments and our code1 are available open source

University of Szeged

Data Augmentation for Machine Translation via Dependency Subtree Swapping

Author: Barta Botond
Lakatos Dorina Petra
Nagy A
Nanys P
Ács Judit
Publication venue: 'SZTE Hungarian Scientific Society of the Silicate Industry'
Publication date: 01/01/2023
Field of study

SZTAKI Publication Repository

HunSum-1: an Abstractive Summarization Dataset for Hungarian

Author: Barta Botond
Lakatos Dorina Petra
Nagy A
Nyist M K
Ács Judit
Publication venue: 'SZTE Hungarian Scientific Society of the Silicate Industry'
Publication date: 01/01/2023
Field of study

SZTAKI Publication Repository

Bírósági határozatok automatikus mondatszegmentálásának hatékonyságmérése

Author: Csányi Gergely
Fülöp Anna
Lakatos Dorina Petra
Megyeri Andrea
Nagy Dániel
Vadász János Pál
Vági Renátó
Üveges István
Publication venue: Universití of Szeged
Publication date: 01/01/2024
Field of study

SZTE Publicatio Repozitórium - SZTE - Repository of Publications

SIGMORPHON 2021 Shared Task on Morphological Reinflection: Generalization Across Languages

Author: Aiton Grant
Ambridge Ben
Ataman Duygu
Ate Yustinus Ghanggo
Barta Botond
Bayyr-ool Aziyana
Bernardy Jean-Philippe
Chodroff Eleanor
Coler Matt
Cotterell Ryan
Ek Adam
El-Khaissi Charbel
Ganieva Sofya
Gasser Michael
Goldman Omer
Habash Nizar
Hatcher Richard J.
Hulden Mans
Ivanova Sardana
Khalifa Salam
Kieraś Witold
Klyachko Elena
Krizhanovsky Andrew
Krizhanovsky Natalia
Kumar Ritesh
Lakatos Dorina
Lane William
Leonard Brian
Liu Zoey
Mielke Sabrina J.
Montoya Samame Jaime Rafael
Nicolai Garett
Nuriah Zahroh
Oncevay Arturo
Pimentel Tiago
Plugaryov Matvey
Ponti Edoardo M.
Prud'hommeaux Emily
Raj Mohit
Ratan Shyam
Ryskina Maria
Salchak Aelita
Salehi Ali
Shcherbakov Andrey
Sheifer Karina
Silva Villegas Gema Celeste
Stoehr Niklas
Straughn Christopher
Suhardijanto Totok
Szolnok Gábor
Tyers Francis M.
Vania Clara
Vylomova Ekaterina
Washington Jonathan
Woliński Marcin
Wu Shijie
Yarowsky David
Ács Judit
Publication venue: The Association for Computational Linguistics
Publication date: 01/08/2021
Field of study

This year's iteration of the SIGMORPHON Shared Task on morphological reinflection focuses on typological diversity and cross-lingual variation of morphosyntactic features. In terms of the task, we enrich UniMorph with new data for 32 languages from 13 language families, with most of them being under-resourced: Kunwinjku, Classical Syriac, Arabic (Modern Standard, Egyptian, Gulf), Hebrew, Amharic, Aymara, Magahi, Braj, Kurdish (Central, Northern, Southern), Polish, Karelian, Livvi, Ludic, Veps, Võro, Evenki, Xibe, Tuvan, Sakha, Turkish, Indonesian, Kodi, Seneca, Asháninka, Yanesha, Chukchi, Itelmen, Eibela. We evaluate six systems on the new data and conduct an extensive error analysis of the systems' predictions. Transformer-based models generally demonstrate superior performance on the majority of languages, achieving >90% accuracy on 65% of them. The languages on which systems yielded low accuracy are mainly under-resourced, with a limited amount of data. Most errors made by the systems are due to allomorphy, honorificity, and form variation. In addition, we observe that systems especially struggle to inflect multiword lemmas. The systems also produce misspelled forms or end up in repetitive loops (e.g., RNN-based models). Finally, we report a large drop in systems' performance on previously unseen lemmas.Peer reviewe

Edinburgh Research Explorer

Helsingin yliopiston digitaalinen arkisto

Characterization of an Aerosol-Based Photobioreactor for Cultivation of Phototrophic Biofilms

Author: Andreas Weber
Dorina Strieth
Johannes Robert
Jonas Kollmen
Judith Stiefelmaier
Kai Muffler
Marianne Volkmar
Michael Lakatos
Roland Ulber
Volkmar Jordan
Publication venue: 'MDPI AG'
Publication date: 01/10/2021
Field of study

Phototrophic biofilms, in particular terrestrial cyanobacteria, offer a variety of biotechnologically interesting products such as natural dyes, antibiotics or dietary supplements. However, phototrophic biofilms are difficult to cultivate in submerged bioreactors. A new generation of biofilm photobioreactors imitates the natural habitat resulting in higher productivity. In this work, an aerosol-based photobioreactor is presented that was characterized for the cultivation of phototrophic biofilms. Experiments and simulation of aerosol distribution showed a uniform aerosol supply to biofilms. Compared to previous prototypes, the growth of the terrestrial cyanobacterium Nostoc sp. could be almost tripled. Different surfaces for biofilm growth were investigated regarding hydrophobicity, contact angle, light- and temperature distribution. Further, the results were successfully simulated. Finally, the growth of Nostoc sp. was investigated on different surfaces and the biofilm thickness was measured noninvasively using optical coherence tomography. It could be shown that the cultivation surface had no influence on biomass production, but did affect biofilm thickness

Multidisciplinary Digital Publishing Institute

Directory of Open Access Journals

PubMed Central